You have a scanned contract, an old letter or a photographed page from a book, and you need the words in it: to edit, to quote, to search or to paste into an email. But the PDF won’t let you select anything, because as far as your computer knows, it’s a picture.
Optical character recognition (OCR) fixes that by reading the letters in the image and turning them into real text. It works remarkably well on good input and struggles on bad input, so most of this guide is about giving it good input.
Try it: OCR PDFRecognize text in scanned PDFs and photos of documents in 16 languages, including English and Bengali, and export it to Word or plain text.
Open tool — freeHow OCR works, briefly
OCR software looks at an image of a page, finds the areas that contain text, splits them into lines and words, and then matches the shapes of the characters against what it knows about a particular language and alphabet. Language knowledge helps a lot: if a shape could be “rn” or “m”, knowing which spelling makes a real word tips the balance.
Two consequences follow. First, anything that makes characters harder to see, such as blur, shadows, skew or low resolution, makes mistakes more likely. Second, telling the software the right language matters, because it uses that to resolve ambiguous shapes.
Getting a scan that OCR can read
The single biggest improvement you can make happens before you open any OCR tool.
With a scanner
- Scan at 300 dpi. It’s the usual recommendation for OCR, and for good reason. Lower resolutions blur small print; much higher just makes huge files.
- Use grayscale or color, not pure black and white. Strict black-and-white mode can break thin strokes and fill in letters.
- Keep the page straight. Line it up with the scanner edge. Small tilts are usually handled, but crooked pages still cost accuracy.
With a phone
- Shoot straight down, with the page filling the frame. Angled shots squash the letters at the far edge.
- Use daylight or even room light. Avoid the flash, which leaves a glare spot, and avoid your own shadow across the page.
- Flatten the page. Curved book pages near the spine are the hardest case. Press the book flat or use a scanning app that corrects curvature.
- Hold still and tap to focus. A slightly blurred photo looks fine to you and terrible to OCR.
If the page is sideways or upside down in the PDF, rotate it first. Our guide on how to rotate and reorder PDF pages covers that.
Choosing the right language
Our OCR tool recognizes text in 16 languages, including English and Bengali. Always pick the language the document is actually written in. Running a French letter through English OCR will mangle accents and produce odd substitutions, and a document in a non-Latin script needs its own language model to work at all.
Mixed-language documents are trickier. If most of the text is in one language with a few words of another, choose the main language and expect to fix the rest by hand. If a document is evenly split, consider running OCR twice and taking each section from the run that matches it.
How to convert a scanned PDF to text step by step
- Check you actually need OCR. Try selecting text in your PDF viewer. If words highlight, the PDF already has text, and a regular PDF to Word conversion is faster.
- Open the OCR tool and add your scanned PDF. Processing happens in your browser, so the document isn’t uploaded. Larger files take a little longer because your own device does the work.
- Select the document’s language.
- Run the recognition and let it work through every page.
- Export the result as a Word file if you want to edit and reformat, or as plain text if you just need the words.
- Proofread against the original, focusing on the error-prone spots listed below.
For a single photo, screenshot or receipt rather than a PDF, our image to text tool does the same job and lets you copy the recognized text straight away. Phones have this built in too: Live Text on recent iPhones and Google Lens on Android are handy for a quick paragraph.
Where OCR goes wrong, and how to check
Even good OCR makes predictable mistakes. Knowing where to look makes proofreading much faster than reading every word.
| Watch for | Typical error |
|---|---|
| Similar shapes | 0 and O, 1 and l and I, 5 and S, rn and m |
| Numbers and amounts | Dropped decimal points, misread digits |
| Names and addresses | Unusual words “corrected” to common ones |
| Multi-column pages | Lines from two columns merged into one |
| Tables | Cells run together as plain lines of text |
| Stamps, signatures, logos | Read as random characters |
For anything where accuracy matters, such as figures in a contract or reference numbers, compare each one against the scan. A spell checker will catch many word errors but won’t notice a wrong digit.
If your goal is tabular data from statements or invoices, OCR to plain text isn’t ideal, since it loses the columns. See our guide to converting bank statement PDFs to Excel instead.
What OCR won’t do well
- Handwriting. Standard OCR is built for print. Neat capitals sometimes work; cursive rarely does.
- Very small or faded text. Photocopies of photocopies lose too much detail.
- Complex layouts. Magazines, forms and brochures come out as text in a sensible order at best, not as a replica of the design.
- Decorative fonts. Script and heavily stylized type are hard to recognize reliably.
In those cases, OCR can still save you typing, but plan on a careful edit afterwards.
A routine that works
Scan at 300 dpi in grayscale, straight and evenly lit. Fix rotation, pick the correct language, export to Word if you’ll edit or plain text if you only need the words, and proofread numbers and names against the original. Spending two extra minutes on the scan usually saves twenty minutes of corrections.
Frequently asked questions
Why can't I copy text from my scanned PDF?
A scanned PDF stores each page as an image, so there are no actual characters to select. OCR software has to recognize the letters in that image and turn them into real text first.
Can OCR read handwriting?
Neat block capitals sometimes work, but cursive and quick notes usually produce poor results with standard OCR. It is designed mainly for printed text.
How accurate is OCR?
On a sharp, straight scan of clearly printed text, OCR is usually very accurate. Accuracy drops with blur, low resolution, unusual fonts, stamps, shadows and the wrong language setting, so always proofread names and numbers.



